This article originally appeared in The Bar Examiner print edition, Summer 2026 (Vol. 95, No. 2), pp. 18–22.

Cover of NextGen Comprehensive Report

In May 2026, NCBE published a comprehensive report on the NextGen UBE, a capstone on years of development, testing, and validation as the new exam’s first administration approaches in July 2026. Excerpts are produced here for the benefit of Bar Examiner readers after a brief introduction. To view the full report, visit ncbex.org/statistics-research.

Introduction

At the time of publication, 10 jurisdictions are preparing to administer the NextGen UBE for the first time. Such a critical moment in the NextGen UBE journey rests on years of work and outreach that has brought NCBE, jurisdictions, candidates, and bar admissions officials to the inaugural launch. A high-level overview is presented in The NextGen Uniform Bar Examination: A Comprehensive Report on Design, Development, and Delivery.

The report revisits prior phases of the development and testing arc (including pilot, field, and prototype testing) before diving into the NextGen UBE’s digitally native ecosystem and empirical evidence gathered during the January 2026 beta administration. The beta exam not only reaffirmed findings from previous testing phases but also evaluated the NextGen UBE system end to end: content, delivery platform, scoring workflows, jurisdictional interfaces, and data capture.

Findings from the beta administration confirm that the NextGen UBE is ready for operational launch.

Part I Summary: Origins and Foundations

Part I of this report establishes the conceptual and structural foundation of the NextGen UBE.

The development process began with a defined licensure purpose: measurement of minimum competence for newly licensed lawyers. That purpose was translated into construct definitions grounded in empirical practice analysis and stakeholder input. The resulting domains of foundational knowledge and skills define the target of measurement.

The Blueprint operationalizes those domains into enforceable specifications governing content coverage, item-type weighting, and section composition. Item-family specifications further constrain how those specifications are translated into observable performance tasks. Together, these layered design controls establish a traceable chain from licensure purpose to exam structure.

By constraining construct definition, blueprint architecture, and form assembly rules, the examination design reduces ­construct-irrelevant variance and supports cross-form comparability. The foundational architecture described in Part I provides the structural basis for psychometric evaluation and score interpretation.

See pages 7–15 of the Comprehensive Report for Part I: Origins and Foundations in full: ncbex.org/statistics-research.

Part II Summary: A Digital Evolution

Part II describes the digital, administrative, and scoring infrastructure through which the NextGen UBE is delivered and evaluated.

The ecosystem integrates candidate readiness, jurisdiction oversight, secure content delivery, structured monitoring, and analytic scoring within defined system boundaries. APIs connect system components while preserving role-based access controls and content security. Administrative and scoring environments are intentionally separated to protect independence of evaluation.

Operational administration is governed by standardized policies, structured monitoring tools, and documented incident workflows. Response-capture and reconciliation processes ensure completeness and auditability of examinee data.

The scoring architecture combines automated selected-response scoring with independent double grading of constructed responses, supported by analytic rubrics, qualification thresholds, reconciliation controls, and continuous oversight monitoring. Audit trails are preserved at each stage.

Collectively, these systems provide the infrastructure necessary to support the scoring, generalization, and comparability inferences underlying the exam’s validity argument.

See pages 16–32 of the Comprehensive Report for Part II: A Digital Evolution in full: ncbex.org/statistics-research.

Part III Summary: Testing and Development

Part III documents the staged empirical program used to develop and confirm the NextGen UBE as an integrated licensure assessment system. Across pilot, field test, prototype, and beta phases, NCBE evaluated item family functioning, structural coherence, scoring reliability, fairness indicators, and operational robustness under progressively more realistic ­delivery conditions.

The pilot phase served as the primary construct-refinement stage. The field test expanded scale and heterogeneity of the participant pool and generated early technical evidence.
The prototype administration represented the first full-form psychometric evaluation. And the beta administration functioned as the operational confirmation phase.

Unlike earlier administrations, beta was designed to evaluate the complete ecosystem end to end—including readiness workflows, live delivery stability, response preservation and recovery, scaled grading operations, and multi-site jurisdictional monitoring—under launch-ready conditions. From a measurement standpoint, the beta provided the first operationally representative dataset with which to replicate prototype psychometric findings, confirm stability of item and score distributions across forms, and evaluate the behavior of the reporting scale and decision region under realistic examinee engagement and administration constraints.

Finally, beta analyses were used to validate—not re-derive—the recommended passing score range framework.

See pages 33–60 of the Comprehensive Report for Part III: Testing and Development in full: ncbex.org/statistics-research.

Additional Resources


The NextGen UBE is operationally ready for launch. Part IV of the Comprehensive Report, which covers this topic, follows in full

Part IV. Operational Readiness

The development of the NextGen UBE was intentionally structured as a staged empirical program. Each phase of the testing arc—pilot testing, field test, prototype exam, and beta administration—served a distinct purpose in evaluating construct representation, measurement stability, scoring reliability, fairness indicators, and delivery integrity. Part IV synthesizes the accumulated evidence to assess whether the examination is ready for operational implementation and to define the framework that will govern post-launch monitoring.

Operational readiness in high-stakes licensure assessment requires more than acceptable reliability coefficients or positive user feedback. It requires convergence across psychometric, operational, technological, and governance domains. The examination must measure the intended construct, classify examinees consistently at decision thresholds, function without delivery distortion, and maintain fairness across demographic groups. It must also operate cohesively as a system rather than as isolated components.

The evidence documented in this report demonstrates that those conditions have been met.

Replication and Stability Across Administrations

A central principle of defensible measurement is replication.

The prototype administration established the NextGen reporting scale, calibrated items under a Rasch-family modeling framework, evaluated dimensional structure, conducted national standard setting, and performed concordance analyses with the legacy UBE. The beta administration was designed not to generate new psychometric theory but to confirm that those findings persist under operationally representative conditions.

The beta results demonstrate that the structural properties observed during prototype remain stable. Item difficulty distributions replicate within expected tolerance bands. Discrimination indices remain within acceptable ranges and show no systematic degradation following blueprint finalization. Reliability estimates continue to meet high-stakes licensure standards. Dimensionality analyses support interpretation of a unified composite score representing minimum competence for entry-level legal practice.

Importantly, the beta administration did not introduce structural distortion in score distributions. Means and standard deviations aligned with projected scale parameters. No evidence of compression at the scale floor or ceiling emerged. Overall measurement precision, including in the decision region corresponding to the recommended passing score range, remained stable. These findings confirm that the reporting scale functions as intended when applied to a fully integrated, operationally realistic administration.

Scoring Integrity at Scale

The NextGen constructed-response scoring ecosystem represents a central innovation of the NextGen UBE.

The transition to analytic rubrics, independent double grading, structured adjudication workflows, and digital scoring platforms required careful evaluation under increasing volume.

The beta administration extended prototype scoring evidence by stress-testing grading workflows at substantially greater scale. Inter-rater agreement remained strong. Adjudication rates were consistent with, and in several cases modestly lower than, those observed during the prototype administration. Score distributions for constructed-response components remained stable across forms. No evidence suggests that scaling introduced additional variability or compromised scoring precision.

Equally important, the grading platform operated with consistent routing logic, audit trails, and reconciliation workflows. These controls are integral to licensure defensibility, as they ensure that scoring decisions are traceable, reviewable, and governed by predefined tolerance thresholds. The beta administration confirms that the scoring model is scalable without loss of integrity.

Delivery Integrity and System Cohesion

High-stakes examinations must preserve response integrity under real-world conditions.

Psyhometric properties are meaningful only if delivery conditions do not introduce ­construct-irrelevant variance.

The beta administration incorporated full-system delivery using the finalized digital ecosystem, including candidate readiness workflows, jurisdiction monitoring interfaces, secure delivery applications, failover recovery mechanisms, and integrated data flows.

Intentional failover simulations confirmed that response preservation protocols function as designed. No response data were lost during simulated interruptions. Completion rates exceeded 99 percent, and no examinee failed to complete due to technology instability. Multi-site administration across jurisdictions did not reveal form-level anomalies or site-specific distortions.

These findings are operationally significant and psychometrically consequential. Stable delivery conditions protect comparability across administrations and ensure that score variation reflects examinee ability rather than technological artifact.

Fairness and Subgroup Stability

Across the testing arc, subgroup performance patterns have remained consistent with historical licensure testing patterns.

Observed subgroup differences in mean item p-values during the beta administration were consistent in magnitude and direction with prototype findings and did not indicate the emergence of new structural disparities.

The replication of fairness indicators under operational conditions strengthens confidence that blueprint stabilization and ecosystem maturation did not introduce unintended measurement distortion.

Validation of the Passing Score Framework

The recommended passing score range was derived from a synthesis of standard-setting judgments, a statistical concordance study, and outcome modeling.

The beta administration provided the first opportunity to evaluate the behavior of that framework under operationally realistic conditions.

Beta findings confirm that the reporting scale remains stable, that overall measurement precision remains within acceptable bounds, and that pass-rate behavior across the 610–620 range is consistent in pattern with prototype results. No evidence of discontinuities in classification outcomes emerged across mapped legacy-equivalent cut scores.

These results provide empirical support for the continued use of the recommended passing score range during early operational administrations, subject to each jurisdiction’s independent authority to establish its own standard.

Risk and Ongoing Oversight

Operational readiness does not imply that the system will remain static.

It requires a defined framework for ongoing monitoring. Early operational administrations will include structured review of distributional stability, inter-rater agreement metrics, adjudication rates, DIF flag rates, and platform performance indicators. These metrics will be evaluated relative to prototype and beta benchmarks to ensure continuity.

Where deviations emerge, the governance structure established during development will support timely review and adjustment. Continuous monitoring is an integral component of responsible licensure testing and will remain embedded within the NextGen operational model.

Conclusion

The NextGen UBE represents a measured evolution of licensure assessment—one grounded in evidence, shaped by stakeholder input, and tested under conditions that reflect real-world administration.

Across multiple phases of development, NCBE has evaluated not only individual components of the exam, but the performance of the full system: content, platform, administration, and scoring. The evidence presented in this report demonstrates that these components function together as intended, producing stable, reliable, and defensible measures of minimum competence.

The psychometric results are consistent and replicable. The operational model has been tested at scale. The digital platform has demonstrated stability, accessibility, and resilience under live conditions. Scoring processes have been implemented with the controls and oversight required for high-stakes decision-making. Importantly, these outcomes have been observed not in isolation, but through an integrated testing arc designed to reduce uncertainty at each stage of development.

Innovation in licensure assessment carries inherent responsibility. New item types, digital delivery, and expanded measurement of skills introduce complexity—but they also provide an opportunity to better align licensure testing with the realities of modern legal practice. Throughout this process, NCBE has approached that responsibility conservatively, ensuring that each advancement is supported by empirical evidence and does not compromise the standards that jurisdictions rely on.

The conclusion is not that the exam is simply different. The conclusion is that it works.

It works as a measurement system. It works operationally. And it works in service of the core purpose of licensure: protecting the public by ensuring that newly licensed lawyers are prepared for safe and effective practice.

The NextGen UBE is ready for operational launch.

Contact us to request a pdf file of the original article as it appeared in the print edition.

  • Bar
    Bar Exam Fundamentals

    Addressing questions from conversations NCBE has had with legal educators about the bar exam.

  • Online
    Online Bar Admission Guide

    Comprehensive information on bar admission requirements in all US jurisdictions.

  • NextGen
    NextGen Bar Exam of the Future

    Visit the NextGen Bar Exam website for the latest news about the bar exam of the future.

  • NCBE
    NCBE Official Study Aids

    Prepare for the bar exam and MPRE using NCBE’s authentic and affordable study aids.

  • 2025
    2025 Year in Review

    NCBE’s annual publication highlights the work of volunteers and staff in fulfilling its mission.

  • 2025
    2025 Statistics

    Bar examination and admission statistics by jurisdiction, and national data for the MBE and MPRE.